Back

Artificial Intelligence in Medicine

Elsevier BV

Preprints posted in the last 30 days, ranked by how well they match Artificial Intelligence in Medicine's content profile, based on 17 papers previously published here. The average preprint has a 0.03% match score for this journal, so anything above that is already an above-average fit.

1
PhenoXtract: combining Large Language Model and Knowledge Graph embedding to extract phenotypes from clinical descriptions

Berardelli, S.; BRIERE, G.; Loire, B.; De Paoli, F.; Gazzo, A. M.; Limongelli, I.; Magni, P.; Zucca, S.; Baudot, A.

2026-06-26 genomics 10.64898/2026.06.22.733382 medRxiv
Top 0.1%
5.5%
Show abstract

Motivation: Standardized phenotypic descriptions are essential for accurate diagnosis, yet clinicians and researchers face challenges in manually extracting and mapping phenotypes from scientific literature or patient clinical records to the Human Phenotype Ontology. Recent advances in deep learning offer new opportunities for automation. We developed PhenoXtract, a novel phenotype extraction approach that combines Large Language Models and Knowledge Graph embedding. PhenoXtract is a multistep pipeline that takes clinical descriptions as input, extracts candidate phenotype entities using large language models, and maps them to terms from an enriched version of the Human Phenotype Ontology, processed as a knowledge graph. Results: Evaluation against expert-curated ground-truth datasets show a recall of 0.70 and precision of 0.85 for PhenoXtract, demonstrating concordance with manually extracted phenotypes, with a computation time of 10-20 seconds for each text analyzed. Moreover, PhenoXtract surpasses rule-based and deep learning-based state-of-the-art tools in two out of the three ground-truth datasets evaluated. These results suggest that hybrid approaches combining Large Language Models and Knowledge Graph embeddings represent a promising direction for automated clinical phenotyping at scale.

2
A multi-view similarity network fusion framework for syndrome discovery from aggregated health records

Gomes Ferreira, A. P.; Anzel, A.; Tavares Veras Florentino, P.; Pereira Ramos, P. I.; Barral-Netto, M.; Marcilio, I.; Hattab, G.

2026-07-02 health informatics 10.64898/2026.06.30.26356975 medRxiv
Top 0.1%
4.3%
Show abstract

Syndrome discovery, the identification of clinically meaningful groupings of signs and symptoms, is a foundational but labor-intensive task in syndromic surveillance, and the COVID-19 pandemic exposed the rigidity of expert-curated definitions in the face of novel threats. Unsupervised, data-driven methods are well-suited to this problem but remain underused. We propose an unsupervised framework based on Similarity Network Fusion (SNF) that operates on only five variables: diagnosis code, sex, age group, epidemiological week and year, and encounter count. Each diagnosis code was represented through three complementary views corresponding to the fundamental questions of syndromic surveillance: what condition is recorded (clinical, via SapBERT embeddings), who is affected (demographic, via chi-square distances), and when it occurs (temporal, via Move-Split-Merge). The fused affinity matrix is partitioned by spectral clustering and exported directly in the Open Syndrome Definition (OSD) format for downstream integration. To our knowledge, this is the first framework of SNF applied to the task of syndrome discovery. We use the framework in 72.9 million primary care encounters across ten Brazilian municipalities. Treating each city as an experiment with no shared training signal yields 47 candidate syndromes, 72% of which are rated fully valid by an expert blind to the procedure. By requiring no predefined targets, the framework discovers candidate syndromes at scale, including ones never explicitly sought, and emits them in a deployable format, shortening the path from emerging signal to usable definition.

3
Automated Interpretation of EEG Reports Using a Large Language Model with Structured Confidence Outputs

Tian, W.; Bergner, S.; Moiseev, A.; Popowich, F.; Medvedev, G.; Richardson, M. P.; Rodionov, R.; Xi, P.; Doesburg, S. M.; Ribary, U.; Winston, J. S.; Vakorin, V. A.

2026-07-10 health informatics 10.64898/2026.07.07.26357190 medRxiv
Top 0.1%
4.3%
Show abstract

Background: Free-text EEG reports typically lack structure, hindering scalable analysis. We evaluate a large language model (LLM) pipeline to extract structured diagnostic labels and confidence levels from these reports. Methods: We developed a hierarchical annotation schema to classify EEG reports for four specific abnormality types using a four-point confidence scale. To establish ground truth, two certified EEG technicians annotated a diverse dataset of reports authored by neurologists with distinct writing styles. We then implemented a grammar-constrained Mistral-7B pipeline, iteratively prompt-tuned on a development set to mirror these expert annotations. The pipeline's effectiveness was evaluated against the human expert benchmark using core agreement (diagnostic accuracy) and certainty-adjusted agreement (confidence alignment), with classical NLP models serving as a secondary baseline. Results: Mistral-7B significantly outperformed baselines, achieving 96% accuracy for overall abnormality detection, approaching the human benchmark of 98%. Crucially, the model successfully identified rare epileptiform abnormalities where traditional models failed and generalized robustly across distinct reporting styles. While diagnostic accuracy was high, a performance gap persisted in certainty-adjusted agreement, indicating that accurately modeling nuanced clinical confidence remains a challenge. Conclusion: LLMs can effectively automate the extraction of structured diagnostic information from EEG reports with near-human accuracy and strong generalization. While confidence calibration requires further refinement, the combination of accurate classification and explainability makes this pipeline a promising tool for standardizing clinical data at scale. Keywords: Routine Clinical Electroencephalography; Large Language Models; Clinical NLP; Confidence Assessment; Explainable AI; Neurophysiological Evaluation

4
Assessment of Zero-Shot Large Language Model (LLM) Assisted Clinical Trial Matching Processes: A Metastatic Cancer Use Case

Weng, Y.; Yalamaddi, H.; Fu, D.; Mishra, A.; Bunning, B. J.; Martin, A. B.; Hope, J.; Charu, V.; Kurian, A.; Desai, M.

2026-07-10 oncology 10.64898/2026.07.06.26354647 medRxiv
Top 0.1%
4.3%
Show abstract

Introduction: For oncology patients with limited treatment options, clinical trials may be a critical lifesaving pathway. Identifying relevant trials, however, is a time-consuming and difficult task. Several patient-trial matching processes incorporating large language models (LLMs) have been proposed to alleviate the burden on patients and oncologists. We aim to explore the benefits and practical challenges of zero-shot LLM-assisted trial matching processes by analyzing the results for a single pancreatic cancer patient. Materials and Methods: The results of a simple zero-shot LLM-assisted clinical trial matching process for our patient were compared to those of a "human benchmark," which was developed manually by two of the authors interfacing directly with ClinicalTrials.gov. Performance metrics -- sensitivity, specificity, precision, and accuracy -- were calculated. In addition, a qualitative content analysis (QCA) of LLM reasoning text was done to identify patterns in "errors," which we define as a human-LLM discrepancy in final patient eligibility. Implications and severity of errors are discussed. Results: The zero-shot LLM-assisted process returned potential trials with a sensitivity, specificity, and precision of 81.1%, 89.3%, and 86.5% respectively compared to the human benchmark. Qualitative error analyses revealed that about 73% of errors could potentially be alleviated with improved prompting and information access. Overall performance seemed comparable to that of human reviewers. Conclusion: The results from this preliminary real-world case study provide additional evidence to the literature in support of the integration of LLMs in clinical trial matching to provide benefit to patients with metastatic cancer with limited options.

5
NEXIM: A Nash Equilibrium-Based Framework for Stable Explainable AI in Medical Applications

Upadhyaya, D. P.; Sahoo, S. S.; Prantzalos, K.; Golnari, P.

2026-07-06 health informatics 10.64898/2026.06.25.26356568 medRxiv
Top 0.2%
3.2%
Show abstract

Reliable explanations are important for trustworthy medical applications of artificial intelligence (AI), but attribution-based explanations can vary across model randomization and small analytic changes. We present NEXIM (Nash Equilibrium-based Explainability and Interpretability Model), implemented here as an accuracy-constrained, equilibrium-inspired model-selection framework that jointly evaluates held-out prediction error, explanation stability, and cross-model connectivity. The implementation evaluated ten GradientBoostingRegressor models per prediction horizon, differing only by random seed (0-9), using a fixed 75/25 patient split. Kernel SHAP attribution vectors were compared using Spearman rank correlation, and graph connectivity summarized whether each model belonged to a dense explanation-similarity region. Candidate models within 0.02 Montreal Cognitive Assessment points of the best root mean squared error (RMSE) were ranked using a multiplicative Explanation Equilibrium Score. In longitudinal Parkinson's Progression Markers Initiative data, NEXIM selected the RMSE-optimal model at the one- and three-year horizons. At the two-year horizon, it selected Model 4 rather than the RMSE-only Model 8, increasing scaled stability from 0.8757 to 0.8847 and normalized graph connectivity from 0.889 to 1.000 while increasing RMSE by only 0.0014. The two models retained the same top-20 feature set but differed modestly in feature order, illustrating that NEXIM primarily acted as a reproducibility screen rather than identifying clinically contradictory explanations. Stability and consensus are treated as reproducibility criteria, not evidence of causal faithfulness, clinical usefulness, or improved patient outcomes. NEXIM may therefore serve as a governance checkpoint for model refresh and documentation, but external validation, stronger model-family baselines, and prospective clinical evaluation remain necessary.

6
Fine-Tuned Large Language Models for Detecting Social Isolation from Unstructured Clinical Notes

Chinthala, L. K.; Lemon, C.; Shaban-Nejad, A.; Farage, G.; Davis, R. L.; Xu, H.; Madlock-Brown, C.

2026-07-07 health informatics 10.64898/2026.07.05.26357334 medRxiv
Top 0.2%
3.2%
Show abstract

Objectives: This study aimed to leverage FLAN-T5-Large, BERT, RoBERTa, and Gemma-2-2B, with fine-tuning, to identify instances of social isolation and social support within unstructured clinical notes. Materials and Methods: Annotated clinical note spans containing social context cues were used to fine-tune each model. Performance was evaluated using Accuracy, Precision, Recall, and Macro-F1 score. A structured prompt was used to instruct the model to perform classification task and mitigate overgeneralization. Performance comparisons across the models assessed sensitivity, robustness, and false positive reduction. Results: FLAN-T5-Large achieved highest performance, with Macro-F1 of 0.92{+/-}0.04, demonstrating balanced results across classes: social isolation (F1 = 0.91{+/-}0.03), no social isolation (F1 = 0.94{+/-}0.05), and social support (F1 = 0.90{+/-}0.04). Gemma-2-2B produced comparable results, with Macro-F1 score of 0.89{+/-}0.10. BERT and RoBERTa achieved lower Macro-F1 scores of 0.77{+/-}0.17 and 0.80{+/-}0.21 respectively, with variability across categories. Discussion: A major contribution of this work is precise identification of multiple concepts related to social connectedness. By integrating annotated examples of both true and false positives, including negations and contextually ambiguous terms, the model better distinguished relevant social context cues from noise. Training on both social isolation and support provided a dual framework for comparative analyses and patient stratification. Conclusion: Transformer-based NLP models, particularly FLAN-T5-Large, demonstrated potential for identifying social isolation and social support in clinical text. These findings support the use of generative AI techniques to enhance detection of social isolation from EHRs, advancing context-aware healthcare analytics.

7
Rare-Class Collapse in ECG-Based Ventricular Tachycardia and Fibrillation Detection: A Systematic Benchmark of Class-Imbalance Mitigation from Reweighting to Cascade Classification

Tiruwa, K. R.

2026-06-29 cardiovascular medicine 10.64898/2026.06.26.26356694 medRxiv
Top 0.2%
2.5%
Show abstract

Ventricular tachycardia (VT) and ventricular fibrillation (VF) are the leading electrical causes of sudden cardiac death, but automated detection is limited by strong class imbalance, where lethal arrhythmias account for fewer than 22% of ECG segments. In this setting, standard classifiers can achieve high accuracy by predicting normal rhythm in most cases while missing many lethal events, a failure mode referred to as rare-class collapse. We evaluated six imbalance-handling approaches: naive logistic regression, inverse-frequency reweighting, label-distribution-aware margin loss (LDAM), cost-sensitive training, two-stage cascade classification, and anomaly detection on 15,614 ECG segments from three PhysioNet databases (VTaC, VFDB, CUDB), with an overall normal-to-lethal ratio of 3.6:1. All methods were assessed at a fixed operating point of 95% specificity using recall, area under the precision-recall curve (AUPRC), and missed-lethal-event rate (MLER). The naive model achieved 45.1% recall (MLER = 0.549), missing 564 of 1,027 lethal events despite 84.1% accuracy. The two-stage cascade performed best, with 65.2% recall, AUPRC of 0.821, and MLER of 0.348, reducing missed events by 37% and achieving the highest decision-curve net benefit. Per-source analysis showed near-complete VF detection (recall up to 0.975) but much lower VT detection (recall 0.183), suggesting a feature-space limitation due to spectral similarity between organized VT and rapid sinus rhythm. Overall, the results show that evaluation metrics strongly influence the visibility of rare-class failure, and that cascade-based methods outperform simpler reweighting approaches for detecting lethal arrhythmias.

8
Dynamic Graph Representation Learning for Data-Driven Huntington's Disease Staging: Evaluation Against Existing Embedding Methods and State-Space Models

Abu Zohair, L. M.; Zantout, H.; Gow, A. J.; Woodward, J.; Lones, M.; Vallejo, M.

2026-06-30 health informatics 10.64898/2026.06.27.26355575 medRxiv
Top 0.2%
2.4%
Show abstract

Huntington's disease (HD) presents a heterogeneous neurodegenerative course, with motor, cognitive, and functional symptoms progressing differently across individuals. This atypical progression complicates the definition of discrete disease stages, hindering understanding of disease trajectories, timely pa- tient care, and therapy development. Consequently, current clinical staging systems rely heavily on clinician-defined, domain-specific criteria and fixed clinical measurement boundaries for stage assignment, reducing objectivity and often leading to overlapping clinical measurements across stages. While machine learning methods can help, existing approaches cannot fully capture complex temporal relationships within and across patients. We propose URL- STFN, a dynamic graph-based representation learning model that encodes both inter- and intra-patient temporal patterns from longitudinal clinical measures. We then evaluate disease stages formed through clustering and stability analysis of URL-STFN latent representations, and compare them with representations obtained from conventional embedding approaches. We further benchmark these clustering-based stages against states derived from conventional temporal models, including DHMM. We hypothesize that clustering URL-STFN latent representations enables identification of HD stages with reduced overlap in clinical measurements. The proposed framework is evaluated using 1,477 clinical visits from the Enroll-HD dataset, a large lon- gitudinal cohort with repeated clinical assessments. For staging, we used 44 clinical measurements spanning motor, cognitive, and functional domains. URL-STFN identifies clinically meaningful HD stages consistent with estab- lished disease progression while reducing overlap in clinical feature values compared with DHMM-derived and clinical staging approaches. These find- ings highlight the potential of a dynamic graph-based representation learning and clustering framework to support more objective, data-driven, and precise HD staging.

9
MedZone Embedder: a framework for representation learning of Japanese secondary medical care areas from a national ICU registry, characterizing intensive care provision structure and regional vulnerability

Ohno, K.; Hashimoto, S.

2026-07-20 health informatics 10.64898/2026.07.17.26358373 medRxiv
Top 0.3%
2.1%
Show abstract

Background: In Japan, acute inpatient care is divided into approximately 335 secondary medical care areas, which serve as the basic units for planning healthcare delivery systems under the 8th National Health Care Plan. While comparisons between regions and facilities typically rely on a single risk-adjusted metric, this approach confuses differences in patient demographics with differences in the actual infrastructure of intensive care units (ICUs). This paper presents a framework - MedZone Embedder - for deriving data-driven indicators of regional structural vulnerability by mapping secondary medical care areas onto a learned similarity space, together with its working implementation. The paper sets out the concept, the method, a proof of concept, and an explicit staged validation program, rather than national empirical results. Methods: Each area is represented by a feature vector consisting of aggregated values of intensive care provision indicators derived directly from the Japan Intensive Care Patient Database (JIPAD) - specifically, risk-adjusted mortality rates (standardized mortality ratios and an in-hospital composite indicator), technical efficiency, length of stay, readmission rates, case severity, and case composition - with the within-area variance of these indicators also taken into account. No hierarchical processing by facility type is performed. A contrastive autoencoder (multilayer perceptron encoder 32 -> 16 -> 8, symmetric decoder) is trained by self-supervised learning, using an objective function that combines reconstruction and normalized temperature cross-entropy (NT-Xent) on noise-augmented views. The resulting 8-dimensional embedding supports area searches based on cosine similarity and anomaly scoring in the embedding space (using isolation forest, Mahalanobis distance, or k-nearest-neighbor density), which is normalized to a vulnerability score ranging from 0 to 1. If deep learning libraries are unavailable, or if the number of areas is small, an alternative method using deterministic principal component analysis is employed. Results: This method was implemented and deployed within an operational ICU decision support system on a managed cloud platform. The proof of concept (PoC) is structured around five secondary medical care areas within Kyoto Prefecture and runs entirely on synthetic facility-level aggregate data constructed to follow the JIPAD indicator schema; no registry data were accessed. It generated: an aggregate provision profile for each area; an area embedding space equipped with a similar-area search function; and a vulnerability ranking that identifies areas with low patient numbers and low diversity that exhibit overall poor outcomes. At this scale, the contrastive autoencoder falls back to principal component projection. The deep learning pathway has been implemented and unit testing has been completed; training and evaluation on actual registry data are pending data-use approval and the expansion of data integration. Validation is staged: Stage 2 will train the contrastive pathway over JIPAD-covered areas to assess construct validity against public structural indicators (ICU/HCU beds, population, accessibility), and Stage 3 will extend coverage to all areas via National Database (NDB) linkage. Conclusion: MedZone Embedder reframes regional comparison from single-indicator ranking to structural representation: which areas are alike, and which are structural outliers. The contribution of this paper is the framework - the proposal that the intensive care provision structure of Japanese secondary medical care areas can be learned from a national outcomes registry and read through the lens of what we call institutional debt - together with a deployed implementation and a pre-specified validation program. To our knowledge, this is a candidate first application of contrastive representation learning to Japanese secondary medical care areas.

10
Enhancing Title and Abstract Priority Screening Through SimEd AI Pipeline.

Lecot, P.; Tonoli-Catez, H.; Noseda, A.; Buisse, T.; Chanut, S.; Chanel, I.

2026-06-30 health informatics 10.64898/2026.06.26.26356718 medRxiv
Top 0.3%
1.7%
Show abstract

The exhaustive identification of evidence is central to systematic reviews, but the screening of titles and abstracts remains particularly labor intensive. Priority screening, an active learning approach that ranks records by estimated relevance, has emerged as an effective strategy to reduce screening workload. Its efficiency is commonly quantified using work saved over sampling at 100% recall (WSS@100%), representing the percentage reduction in effort compared with random screening. Although modern priority-screening models achieve high efficiency on many benchmark datasets, some reviews still exhibit low WSS@100%, indicating suboptimal retrieval. Our study sought to improve the retrieval of all relevant articles in challenging datasets to ensure better generalization of priority screening. We first showed using SYNERGY benchmark datasets that while the most advanced ELAS_h3 priority screening model from state-of-the-art ASReview LAB v.2 open-source software, efficiently retrieved most relevant articles, it struggled with the rare, final ones in challenging datasets. To address this, we tested a hybrid approach entitled SimEd AI: using ELAS_h3 for early retrieval and then applying supervised fine-tuning to the biomedical transformer BioMed-RoBERTa-base with these relevant articles to enhance the detection of the remaining difficult cases. We found that fine-tuning BioMed-RoBERTa-base model with 10 late-identified relevant and 10 hard irrelevant study titles and abstracts, enabled faster retrieval of articles of interest compared to ELAS_h3 alone. This approach increased WSS@100% from 46.5% (SD0.0%) to 83.3% (SD0.4%), while adding only an average of 22 minutes of computational time for fine-tuning and inference. The SimEd AI priority screening pipeline could be valuable for situations requiring highest possible recall. It could be particularly useful in scoping reviews with broad or diverse topics where traditional priority screening methods may miss subtle relevance signals. Further work should define a data-driven stopping rule for ending screening once the fine-tuned domain-specific transformer is applied at the final stage and assess generalizability across additional challenging datasets.

11
Beyond Single Biomarkers: A Graph Neural Network Framework for Multivariable Prediction of Clinical Outcomes from Brain Imaging

Esmaelpoor, J.; Kadkhodamohammadi, A.; Peng, T.; Jelfs, B.; Mao, D.; Ghafouri, A.; Shader, M.

2026-06-24 health informatics 10.64898/2026.06.21.26356202 medRxiv
Top 0.3%
1.7%
Show abstract

Understanding brain-behavior relationships requires models capturing the distributed, interactive, and multiscale nature of neural systems. Traditional univariate approaches and single-biomarker models are inherently limited in this context, as they fail to represent dependencies across regions and the hierarchical organization of brain networks. In this study, we propose a graph-based multivariable framework for brain imaging analysis that integrates key organizational principles of brain function-including segregation, integration, modularity, and temporal dynamics-within a unified graph neural network architecture. The framework represents brain data as hierarchical graphs, where node features encode regional activation and temporal variability, and graph structure captures interactions within and between functional modules. The proposed approach is evaluated using functional near-infrared spectroscopy (fNIRS) data as a case study, where subject-specific brain graphs are constructed from task-based recordings acquired shortly after cochlear implant activation to predict speech understanding outcomes one year later. Under leave-one-subject-out validation, the model demonstrates strong predictive performance (R = 0.73, p < 0.001), outperforming previously reported single-biomarker approaches. Perturbation-based analyses further show that predictions are driven by distributed patterns of activity and interaction across regions and modalities, rather than isolated features. These results illustrate the capability of the proposed framework to capture complex brain organization and highlight its potential as a generalizable platform for multivariable analysis and prediction in neuroimaging applications beyond the specific clinical use case considered here.

12
HGGT:Heterogeneous Gated Graph Transformer for Predicting Clinical Trial Success

Qian, L.; Lu, X.; Haris, P.; Yang, Y.

2026-07-01 health informatics 10.64898/2026.06.28.26356795 medRxiv
Top 0.4%
1.7%
Show abstract

Clinical trials are critical milestones in the drug development pipeline, yet their high failure rates and substantial costs underscore the need for robust predictive models. This study introduces a Heterogeneous Gated Graph Transformer (HGGT) model tailored to predict clinical trial success. Unlike existing methods that typically model trial-related entities in isolation or with homogeneous graphs, HGGT explicitly models the rich heterogeneous relationships among trials, diseases, drugs, genes, targets, abstracts, and eligibility criteria through a gated graph transformer architecture, which dynamically learns and weights multi-type relational interactions to capture complex biological and clinical dependencies. By integrating heterogeneous graph representation with transformer-based context modeling, HGGT effectively captures non-linear, multi-scale interactions across biomedical entities, leading to improved predictive performance for trial success. Experimental results demonstrate that the HGGT model achieves strong performance, with the highest PR-AUC, F1 score, and ROC-AUC across three phases. These findings highlight the potential of graph-based deep learning approaches in optimizing clinical trial design and resource allocation, ultimately accelerating the translation of novel therapies into clinical practice.

13
Revealing Hidden Myocardial Infarction Signatures from Brief Single-Lead Electrocardiograms: A Novel Framework for Smart Wearable Applications

Alavi, R.; Li, J.; Matthews, R. V.; Pahlevan, N. M.; Kloner, R. A.; Gharib, M.

2026-07-13 cardiovascular medicine 10.64898/2026.07.08.26357521 medRxiv
Top 0.4%
1.7%
Show abstract

The electrocardiogram (ECG) contains rich nonlinear and non-stationary dynamic information that is only partly captured by conventional ECG interpretation and beat-to-beat metrics, and is increasingly analyzed using black-box artificial intelligence models that often lack interpretability. Here, we introduce the ECG time-frequency "eyeball", an interpretable framework that transforms a brief single-lead ECG recording into a geometric signature and a set of low-dimensional rotational and geometrical features using empirical mode decomposition and Hilbert-based analytic signal mapping. In 30-second lead I-equivalent recordings from 170 healthy subjects and 80 patients with acute myocardial infarction (AMI), the proposed "eyeball" metrics significantly differentiated groups, with AMI associated with higher rotational frequency metrics, lower envelope metrics, and displaced centroid location. Representative examples revealed a coherent morphologic spectrum from normal patterns to geometries consistent with myocardial ischemia, injury, and infarction. The representation remained stable across recording windows from 30 seconds to 5 minutes, and individual "eyeball" features achieved areas under the receiver operating characteristic curve (AUCs) of up to 0.78 for AMI detection. These findings suggest that the ECG time-frequency "eyeball" condenses clinically relevant nonlinear ECG dynamics into an interpretable representation that may reveal hidden AMI signatures, complement conventional ECG interpretation, and provide a foundation for accessible single-lead cardiovascular screening using future smart wearables.

14
FHIRBench: Benchmarking FHIR Clinical Data Serialization Strategies for Large Language Models

Chong, J.

2026-07-15 health informatics 10.64898/2026.07.14.26358020 medRxiv
Top 0.4%
1.5%
Show abstract

We present FHIRBench, a benchmark evaluating six FHIR clinical data serialization strategies across four frontier LLMs (Claude Sonnet 4.5, GPT-5.4, DeepSeek V3.2, Qwen3 32B) on three clinical tasks using 100 stratified synthetic FHIR R4 patient bundles. We employ two evaluation layers: token-level F1 and LLM-as-judge rubric on four clinical dimensions, yielding 7,200 evaluations per layer. Our findings reveal four results. First, serialization significantly impacts quality but the direction diverges between layers: Condensed outperforms Raw JSON on F1 for 3/4 models (Wilcoxon p < 10^-17), while Raw JSON achieves higher judge scores for 3/4 models (p < 10^-7). Narrative achieves 95% of Raw JSON's quality at 83% fewer tokens. Second, model rankings completely reverse between layers -- Claude ranks last on F1 but first on clinical quality (p = 1.0 x 10^-6), demonstrating that single-metric evaluation produces misleading model selection. Third, a significant Model x Serializer interaction (Friedman p = 0.0009) precludes universal format recommendations, with GPT-5.4 favoring Raw JSON while open-weight models favor compressed formats. Fourth, Llama 3.1 70B exhibits 100% inference failure on complex patients despite operating within its nominal context window, revealing a patient-safety gap where AI fails for the patients who need it most. These findings establish that clinical AI systems require model-aware serialization middleware, multi-layer evaluation frameworks, and capacity verification before deployment. Code and data publicly available.

15
MIRA-Net: A Cross-Cohort Representation Learning Framework for Parkinson's Disease Classification Using Acoustic and Beta-Band MEG Biomarkers

Akhila, N.; Ekbal, A.; Roy, D.

2026-07-06 health informatics 10.64898/2026.07.03.26357258 medRxiv
Top 0.4%
1.5%
Show abstract

Accurate diagnosis of Parkinson's disease (PD) remains challenging due to substantial inter-subject variability and the absence of widely accessible, objective multimodal biomarkers. Although speech and magnetoencephalography (MEG) biomarkers have individually demonstrated strong discriminative potential, their joint utilization is constrained by the absence of subject-level paired datasets - a fundamental gap that has prevented cross-modal validation at the individual level. We argue that this makes cross-cohort representation learning not merely a pragmatic workaround, but the most realistic and clinically transferable framework for multimodal PD assessment. In real-world deployment, acoustic screening and neuroimaging biomarkers are acquired through separate clinical pathways and must be integrated across heterogeneous patient populations. To address this, we propose MIRA-Net (Modality-Invariant Residual Adversarial Network). This cross-cohort representation learning framework integrates acoustic speech features from four established UCI datasets (n = 193) with beta-band MEG biomarkers from the NatMEG-PD dataset (n = 127) for PD classification. MIRA-Net employs RF-SHAP feature selection, gradient-reversal-based domain adaptation, and supervised contrastive alignment to learn participant-independent, modality-invariant embeddings. The framework is evaluated under Rest, Go, and Passive task conditions against Early Fusion, Vanilla DANN, and Supervised Contrastive Learning baselines. MIRA-Net achieves a peak accuracy of 86.23% (Go condition, Stacking classifier) with AUC values exceeding 0.88 under repeated cross-validation, alongside a sensitivity of 89.4% and specificity of 83.1%. Friedman tests confirm statistically significant performance differences among fusion strategies (p < 0.003 across all conditions). These results demonstrate that cross-cohort representation learning can extract robust disease-discriminative signatures without synchronized multimodal recordings, offering a practical pathway toward AI-assisted PD assessment in resource-constrained clinical settings.

16
Exploring the Application of the Observational Medical Outcomes Partnership Common Data Model to Multi-site Stroke Rehabilitation Research Data

Loomis, K. J.; Kumar, A.; Marin-Pardo, O.; Bellinger, G. C.; French, M. A.; Roemmich, R. T.; Liew, S.-L.

2026-07-08 health informatics 10.64898/2026.06.28.26356618 medRxiv
Top 0.4%
1.4%
Show abstract

Background: Emerging artificial intelligence and machine learning (AI/ML) tools can help generate robust knowledge to support precision rehabilitation approaches for varied patient populations. There is a large amount of research-generated and clinical rehabilitation data available for this purpose; however, a pronounced lack of interoperability prevents large-scale data aggregation. Common data models (CDMs) such as Observational Medical Outcomes Partnership (OMOP) have improved data interoperability across healthcare settings, and more recently, for clinical rehabilitation data, specifically. However, the application of these CDMs to research-generated data has not yet been explored. Therefore, as a foundational step, our study evaluated the breadth and depth of OMOP CDM coverage for data in a multi-site repository of harmonized rehabilitation research data: the Enhancing NeuroImaging Genetics through Meta-Analysis Stroke Recovery (ENIGMA-SR) database. Methods: Two raters independently mapped data elements representing 46 demographics and medical history (DMH) ENIGMA-SR variables and 95 distinct ENIGMA-SR rehabilitation assessments to OMOP standard concepts. Initial rater agreement was assessed for data element inclusion in OMOP and for specific OMOP concepts used (primary metric: Gwet's agreement coefficient [AC]). Mapping differences were reconciled, and final mappings were descriptively analyzed to examine (1) overall OMOP inclusion, (2) inclusion of more granular levels (subscales, items) of complex assessments, and (3) mapped OMOP concept characteristics. Results: Initial rater agreement was good/very good for overall OMOP inclusion of DMH and assessment data elements and for OMOP concepts mapped across almost all assessment data elements (Gwet's AC: 0.79-0.89). Initial OMOP concept agreement was more variable for DMH data elements; however, all mapping differences were successfully reconciled to 100%. Overall, DMH data elements had higher OMOP inclusion than rehabilitation assessments: 84.8% (39/46) vs. 58.9% (56/95). OMOP coverage was particularly limited for complex assessment subscale- and item-level data elements (9.4% [3/32]; 19.2% [14/73]) and did not match the granularity level represented in ENIGMA-SR data for 56.2% (41/73) of complex assessments. DMH and top-level assessment data elements were frequently mapped to multiple OMOP concepts (median: 6, 2; range: 1-23, 1-8), and for > 50% of these data elements the concepts spanned 2-3 different OMOP domains. Conclusion: For ENIGMA-SR, the OMOP CDM has good coverage of DMH data, moderate top-level coverage of rehabilitation assessments, and very limited coverage of assessment subscales and items. This uneven coverage, combined with variability in OMOP concepts and domains mapped to equivalent data points, presents challenges for aggregating clinical and research-generated rehabilitation data into AI/ML-ready datasets. Moreover, software tools currently available to facilitate the mapping process do not effectively accommodate content- and structure-related features inherent to research-generated data. Going forward, the utility of the OMOP CDM to aggregate multi-source rehabilitation data may be improved by expanding the catalogue of OMOP rehabilitation-related concepts, building cross-walks to research-oriented data standards, and adapting emerging computational tools to streamline the mapping process.

17
Large language models for cancer registry abstraction: a real-world evaluation across models, variables, and cancer types

Fuchs, J.; Satusky, M. J.; Leese, P. J.; Nag, S.; Zipple, I. W.; Baggett, C. D.; Lash, S.; Reeder-Hayes, K.; Wood, W. A.; Johnson, C. T.; Critchley, C.; Krishnamurthy, A. K.; Elston Lafata, J.; Thompson, C. A.; Troester, M. A.; Pfaff, E. R.

2026-06-29 health informatics 10.64898/2026.06.25.26356626 medRxiv
Top 0.4%
1.3%
Show abstract

Cancer registries enable cancer surveillance at the population level. These registries require significant human-time to read through many different parts of the electronic health record, including structured data and lengthy, free-text clinical reports, to abstract values for hundreds of required variables. Large language models (LLMs) offer the possibility to significantly improve this process by supporting and speeding up cancer registry data abstraction. However, it is unclear how well these models perform at real-world cancer registry abstraction involving multiple cancer types and large patient volumes. Here, we evaluate five foundational LLMs for their ability to reliably abstract cancer registry variables. We leverage hospital cancer registry data from a large regional health system as the ground truth and use LLMs to abstract from clinical reports eight registry variables for 5,939 patients with seven different cancer types. We use a zero-shot prompting strategy to compare LLM ability on commonly abstracted cancer variables with different data types. The results show that larger and more advanced models (Claude Sonnet 4.5, GPT-OSS-120b, GPT-OSS-20b) generally outperform smaller models (Gemma 12b, LLaMA 3.1 8b). The best performing models show F1 scores around 0.8 for cancer registry variables with low cardinality (grade, summary stage, laterality), with only slightly lower F1 scores for variables with high cardinality (primary site, regional nodes examined, regional nodes positive). On the more complex task of precise date extraction, all models showed decreased performance on both diagnosis and treatment dates (exact accuracy ~0.55 for the best performing models), which increased to ~0.85 for a tolerance within {+/-}30 days. These results quantify the performance of various models as well as the potential and limitations of LLMs in cancer registry abstraction tasks.

18
Image-based deep learning for emergency electrocardiogram classification

Meneguitti Dias, F.; Ribeiro, E.; Olivetti, N.; Carvalho, O.; Krieger, J. E.; Gutierrez, M.

2026-06-22 cardiovascular medicine 10.64898/2026.06.18.26355968 medRxiv
Top 0.5%
1.2%
Show abstract

Automated electrocardiogram analysis has advanced largely through digital waveforms, yet many emergency-care workflows rely on ECGs available only as printed tracings, scanned reports, PDFs or mobile photographs. We developed an image-based deep learning system for emergency ECG classification and evaluated it in InCor-EMG, an expert-adjudicated dataset of 18,519 emergency ECGs spanning 12 ECG categories, with labels from 19 cardiologists. On the held-out test set, the final ConvNeXt ensemble achieved a macro F1-score of 0.807 (95% CI, 0.788-0.825), compared with 0.820 (95% CI, 0.805-0.832) for annotating cardiologists, and higher F1-scores than Mortara Veritas in most evaluated categories. Performance was associated more strongly with inter-reader agreement than with training sample size and remained informative across scanned and photographed ECGs, with supportive performance in model-enriched temporal and heterogeneous public-image evaluations. These findings support ECG image classification when digital waveforms are unavailable.

19
ASTAR: Automated Induction of Standardized Radiology Reporting Templates from Large-Scale Clinical Free-Text Corpora

Zhang, X.; Liu, M.; Chen, Y.; Zhu, J.; Anmahapong, K.; Huang, Y.; Zhang, Y.; Yang, H.; Liao, Y.; Ning, G.; Qu, H.; Tian, Q.

2026-07-14 health informatics 10.64898/2026.07.11.26357801 medRxiv
Top 0.5%
1.1%
Show abstract

Structured reporting converts free-text radiology narratives into queryable data keys, facilitating cohort assembly, longitudinal tracking, and training label generation for medical AI. The prevailing paradigm follows a two-stage pipeline: (1) constructing a reporting template, (2) extracting information to populate it. While the extraction stage has benefited from advances in large language models (LLMs), template construction remains a manual bottleneck relying on labor-intensive expert consensus that is static, difficult to scale, and may fail to capture real-world reporting diversity. We address this limitation with ASTAR, an LLM-based framework for Automated induction of STAndardized radiology Reporting templates from large-scale clinical free-text corpora. Extensive experiments on 4,215 fetal brain MRI reports from multiple centers demonstrate that, in this reporting scenario, the ASTAR-induced template surpasses two expert-curated templates across template coverage, information fidelity, diagnostic fidelity, and expert-rated usability, reducing template development from weeks of committee deliberation to hours of automated processing.

20
Uncertainty-aware extraction of clinical findings from Finnish EHRs using open large language models

Leinonen, J. V.; Knuutila, J.; Kurki, S.; Pamilo, S.; Koskinen, M.

2026-07-09 health informatics 10.64898/2026.07.07.26355248 medRxiv
Top 0.5%
1.1%
Show abstract

Objective. To evaluate whether open-weight large language models (LLMs) can accurately extract clinical findings from Finnish-language pediatric records, and whether prediction uncertainty can be used to triage cases for expert review to minimize manual work. Materials and Methods. Retrospective cohort of 97 pediatric ischaemic stroke patients (1 month - 17 years) from Helsinki University Hospital (2010 - 2023). Three open LLMs (gpt-oss-20b, DeepSeek-R1-Distill-Qwen-32B, and medgemma-27b-text-it) were prompted in English to detect four extraction targets (hemiplegia, headache, seizure, and stroke as a positive control) from each patient's full free-text record. Each combination received 15 calls (five temperatures x three repeats). Performance was benchmarked against a clinician reference (accuracy, recall, precision, F1). Shannon entropy across the 15 calls quantified within-model uncertainty; inter-model disagreement provided an ensemble signal. Patients were ranked by uncertainty for a simulated selective-review workflow. Findings were externally validated in an independent neonatal stroke cohort (n = 88). Results. Gpt-oss-20b achieved the best balance of recall (0.91 - 1.00) and precision (0.83 - 0.92), with F1 0.89 - 0.95 across non-control extraction targets. Entropy in misclassified cases was 2.4 - 3.4 times higher than in correctly classified cases. Entropy-based triage achieved complete error coverage by reviewing <10% of patients for hemiplegia (8.3%) and headache (8.2%), and 19.6% for seizure. Neonatal validation reached F1 0.95 for Apgar 1 min and binary seizure, and F1 0.87 for 4-class stroke-subtype classification. Discussion. Within-model entropy and inter-model disagreement provided complementary, calibrated signals of likely error in a non-English clinical setting. Conclusion. Open LLMs can extract clinical findings from Finnish pediatric records with accuracy comparable to published English benchmarks, and uncertainty-based triage substantially reduces required expert workload.